Papers by Natália da Silva Perez

3 papers
Multilingual Event Extraction from Historical Newspaper Adverts (2023.acl-long)

Copied to clipboard

Challenge: Developing NLP methods for historical corpora is difficult, as only domain experts can label them . off-the-shelf models are trained on modern language texts, rendering them weaker for historical documents .
Approach: They propose to use an annotated newspaper dataset to extract historical data from a novel domain of texts.
Outcome: The proposed method performs well on a multilingual dataset in English, French, and Dutch . it is possible to extract surprisingly good results even with scarce annotated data using existing models and datasets for modern languages .
Measuring Intersectional Biases in Historical Documents (2023.findings-acl)

Copied to clipboard

Challenge: digitised historical documents suffer from errors introduced by optical character recognition (OCR) and are written in an archaic language.
Approach: They investigate the continuities and transformations of bias in Caribbean historical newspapers during the colonial era . they use distributional semantics models and word embeddings to measure gender, race, and intersectional biases.
Outcome: The authors show that gender and racial biases are interdependent and their intersection triggers distinct effects.
A Dual-View Analysis of Multiple Languages in Colonial Newspapers (2026.findings-acl)

Copied to clipboard

Challenge: Historical newspapers from the colonial period offer valuable evidence of how racializing language evolved over time.
Approach: They propose a contextual question answering and visual question answering task from colonial newspapers . they propose linguistic training for temporal word embedding with a compass to study racialization .
Outcome: The proposed tasks are limited for low-resource tasks, the authors show . the authors compare the results of two QA pairs from colonial newspapers to a compass .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations